Papers with manual analysis of annotator
Improving Adversarial Data Collection by Supporting Annotators: Lessons from GAHD, a German Hate Speech Dataset (2024.naacl-long)
Copied to clipboard
| Challenge: | Hate speech detection models are only as good as the data they are trained on, but adversarial datasets are slow and costly . data sourced from social media suffer from systematic gaps and biases, leading to unreliable models with simplistic decision boundaries. |
| Approach: | They propose a German Adversarial Hate speech Dataset comprising 11k examples . they explore new strategies for supporting annotators and provide manual analysis of disagreements for each strategy . |
| Outcome: | The proposed dataset is challenging even for state-of-the-art hate speech detection models and it significantly improves model robustness. |